Papers by Jian Gang Ngui
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages? (2026.acl-long)
Copied to clipboard
Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng, Jann Railey Montalan, Thura Aung, Jian Gang Ngui, Yosephine Susanto, William Chandra Tjhi, Panuthep Tasawong, Erik Cambria, Ekapol Chuangsuwanich, Sarana Nutanong
| Challenge: | SEA-BED examines how multilingual text embeddings perform across tasks and languages . performance gaps arise from data coverage, training objectives, and architectural design, authors say . |
| Approach: | They propose a large-scale benchmark covering 10 SEA languages and diverse embedding tasks. |
| Outcome: | The proposed model performs poorly across languages and tasks, but language-task analyses reveal inconsistencies . the results suggest that performance gaps arise from limitations in data coverage, training objectives, and architectural design. |
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing safeguard models rely on translation of English datasets, missing regional and cultural nuances. |
| Approach: | They propose a framework to generate culturally grounded safety datasets for Southeast Asia . SEA-Guard family is the first multilingual safeguard model grounded in SEA cultural contexts . |
| Outcome: | The proposed model outperforms existing safeguard models in detecting regionally sensitive content while maintaining strong general safety performance. |
SEA-HELM: Southeast Asian Holistic Evaluation of Language Models (2025.findings-acl)
Copied to clipboard
Yosephine Susanto, Adithya Venkatadri Hulagadri, Jann Railey Montalan, Jian Gang Ngui, Xianbin Yong, Wei Qi Leong, Hamsawardhini Rengarajan, Peerat Limkonchotiwat, Yifan Mai, William Chandra Tjhi
| Challenge: | Existing LLM benchmarks are capable of evaluating specific capabilities in English as well as in various mid- to low-resource languages, but a comprehensive and culturally representative evaluation suite for the SEA languages has not been developed thus far. |
| Approach: | They propose a holistic linguistic and cultural LLM evaluation suite that emphasizes SEA languages and introduces a leaderboard that allows users to understand models’ multilingual and multicultural performance. |
| Outcome: | The proposed evaluation suite emphasizes SEA languages and supports Filipino, Indonesian, Tamil, Thai, and Vietnamese. |
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation (2025.acl-long)
Copied to clipboard
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker
| Challenge: | Reliable multilingual evaluation is difficult and culturally appropriate evaluation is even harder to achieve. |
| Approach: | They propose a multilingual evaluation framework that aims to mitigate these biases by improving translations and annotation practices. |
| Outcome: | The proposed framework improves translation quality and cultural coverage and is culturally sensitive and culturally agnostic. |
SEA-SafeguardBench: Culturally Grounded Safety Benchmark for Southeast Asian Languages (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing multilingual safety benchmarks rely on machine-translated English data, which fails to capture nuances in low-resource languages. |
| Approach: | They propose to use a human-verified safety benchmark for Southeast Asian languages to validate their safety and cultural diversity. |
| Outcome: | The proposed model outperforms existing models in general, in-the-wild, and content generation across eight languages and 21,640 samples across three subsets: general, and in- the-wild. |